Papers by Yeo Wei Jie

6 papers
How Interpretable are Reasoning Explanations from Prompting Large Language Models? (2024.findings-naacl)

Copied to clipboard

Challenge: Prompt Engineering has garnered significant attention for enhancing the performance of large language models across a multitude of tasks.
Approach: They propose a simple prompting technique that yields more than 70% improvement in interpretability.
Outcome: The proposed method improves interpretability by 70% across multiple dimensions.
Understanding Refusal in Language Models with Sparse Autoencoders (2025.findings-emnlp)

Copied to clipboard

Challenge: a study of refusal in instruction-tuned language models identifies latent features that causally mediate refusal behaviors.
Approach: They conduct a mechanistic study of refusal in instruction-tuned LLMs using sparse autoencoders . they identify latent features that causally mediate refusal behaviors using sparsed autoencoding .
Outcome: The proposed method validates refusal-related features across multiple datasets.
Plausible Extractive Rationalization through Semi-Supervised Entailment Signal (2024.findings-acl)

Copied to clipboard

Challenge: Abstract: Large language models are gaining widespread adoption in natural language processing tasks.
Approach: They propose a semi-supervised approach to optimize for plausibility of extracted rationales by using a pre-trained natural language inference model and a supervised NLI predictor.
Outcome: The proposed model outperforms unsupervised models by > 100% on a ERASER dataset.
Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are capable of generating persuasive Natural Language Explanations (NLEs) however, the faithfulness of these explanations should not be readily trusted at face value.
Approach: They propose to use a causal mediation technique called activation patching to measure the faithfulness of an explanation towards supporting the explained answer.
Outcome: The proposed metric, Causal Faithfulness, quantifies the consistency of causal attributions between explanations and the corresponding model outputs as the indicator of faithfulness.
Self-training Large Language Models through Knowledge Detection (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) often require extensive labeled datasets and training compute to achieve impressive performance across downstream tasks.
Approach: They propose a self-training paradigm where the LLM curates its own labels and selectively trains on unknown data samples identified through a reference-free consistency method.
Outcome: The proposed model reduces the dependency on large labeled datasets and mitigates catastrophic forgetting in out-of-distribution benchmarks.
SusGen-GPT: A Data-Centric LLM for Financial NLP and Sustainability Report Generation (2025.findings-naacl)

Copied to clipboard

Challenge: Existing tools for financial reporting and ESG analysis are lacking . large language models are not proficient across general finance and ESE domains .
Approach: They propose a dataset that includes seven financial NLP tasks and a benchmark to improve sustainability report generation.
Outcome: SusGen-30k, a high-quality dataset, shows state-of-the-art performance . it surpasses all other models except GPT-4 in six adapted tasks and two off-the shelf tasks .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations